A reliable sentiment analysis for classification of tweets in social networks

Masoud AminiMotlagh; HadiShahriar Shahhoseini; Nina Fatehi

doi:10.1007/s13278-022-00998-2

A reliable sentiment analysis for classification of tweets in social networks

Soc Netw Anal Min. 2023;13(1):7. doi: 10.1007/s13278-022-00998-2. Epub 2022 Dec 12.

Authors

Masoud AminiMotlagh¹, HadiShahriar Shahhoseini¹, Nina Fatehi²

Affiliations

¹ School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran.
² Department of Electrical and Computer Engineering, Wayne State University, Detroit, USA.

Abstract

In modern society, the use of social networks is more than ever and they have become the most popular medium for daily communications. Twitter is a social network where users are able to share their daily emotions and opinions with tweets. Sentiment analysis is a method to identify these emotions and determine whether a text is positive, negative, or neutral. In this article, we apply four widely used data mining classifiers, namely K-nearest neighbor, decision tree, support vector machine, and naive Bayes, to analyze the sentiment of the tweets. The analysis is performed on two datasets: first, a dataset with two classes (positive and negative) and then a three-class dataset (positive, negative and neutral). Furthermore, we utilize two ensemble methods to decrease variance and bias of the learning algorithms and subsequently increase the reliability. Also, we have divided the dataset into two parts: training set and testing set with different percentages of data to show the best train-test split ratio. Our results show that support vector machine demonstrates better outcomes compared to other algorithms, showing an improvement of 3.53% on dataset with two-class data and 7.41% on dataset with three-class data in accuracy rate compared to other algorithms. The experiments show that the accuracy of single classifiers slightly outperforms that of ensemble methods; however, they propose more reliable learning models. Results also demonstrate that using 50% of the dataset as training data has almost the same results as 70%, while using tenfold cross-validation can reach better results.

Keywords: Data mining; Sentiment analysis; Social networks analysis; Text mining.

© The Author(s), under exclusive licence to Springer-Verlag GmbH Austria, part of Springer Nature 2022, Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.